congressional speech
US congressional speeches are getting less evidence-based over time
The language that elected members of the US Congress use in debate increasingly includes words such as "phony" and "doubt" over words such as "proof" and "reason". This linguistic trend away from evidence in favour of intuition was revealed in an artificial intelligence analysis of millions of congressional speech transcripts. It also coincides with both greater political polarisation in Congress and a decline in the number of laws that get enacted through Congress, says Stephan Lewandowsky at the University of Bristol in the UK. How does ChatGPT work and do AI-powered chatbots "think" like us? "We can think that truth is something we can achieve based on analysis of evidence, or we can think of it as the result of intuition or'gut feeling'," says Lewandowsky. "Those notions of honesty and truth are expressed in how we use everyday language."
Structured Embedding Models for Grouped Data
Maja Rudolph, Francisco Ruiz, Susan Athey, David Blei
We study how the word usage of U.S. Congressional speeches varies across states and party affiliation, how words are used differently across sections of the ArXiv, and how the copurchase patterns of groceries can vary across seasons. Key to the success of our method is that the groups share statistical information. We develop two sharing strategies: hierarchical modeling and amortization. We demonstrate the benefits of this approach in empirical studies of speeches, abstracts, and shopping baskets.
A Machine Learning Pipeline to Examine Political Bias with Congressional Speeches
Machine learning, with advancements in natural language processing and deep learning, has been actively used in studying political bias on social media. But the key challenge to model political bias is the requirement of human effort to label the seed social media posts to train machine learning models. Although very effective, this approach has disadvantages in the time-consuming data labeling process and the cost to label significant data for machine learning models is significantly higher. The web offers invaluable data on political bias starting from biased news media outlets publishing articles on socio-political issues to biased user discussions about several topics in multiple social forums. In this work, we introduce a novel approach to label political bias for social media posts directly from US congressional speeches without any human intervention for downstream machine learning models.
Speeches of US politicians 'have the reading age of a 13-year-old'
Congressional speeches made by US politicians have become simpler since the 1970s and only require the reading age of a 13-year-old to be followed, study found. Computer scientists from Kansas State University analysed two million congressional speeches from Republican and Democrat politicians made between 1873 and 2010. Text analysis algorithms were used to examine how congressional speeches changed in terms of complexity, emotion and divisiveness over 138 years. More recent speeches use a smaller vocabulary, simpler language and talk about'the other party' more than speeches made even a decade ago, the authors found. Researchers put the drop in the reading level down to the rise of broadcast media in congress that started in the mid-1970s - with politicians'playing to the camera'.
Structured Embedding Models for Grouped Data
Rudolph, Maja, Ruiz, Francisco, Athey, Susan, Blei, David
Word embeddings are a powerful approach for analyzing language, and exponential family embeddings (EFE) extend them to other types of data. Here we develop structured exponential family embeddings (S-EFE), a method for discovering embeddings that vary across related groups of data. We study how the word usage of U.S. Congressional speeches varies across states and party affiliation, how words are used differently across sections of the ArXiv, and how the co-purchase patterns of groceries can vary across seasons. Key to the success of our method is that the groups share statistical information. We develop two sharing strategies: hierarchical modeling and amortization. We demonstrate the benefits of this approach in empirical studies of speeches, abstracts, and shopping baskets. We show how SEFE enables group-specific interpretation of word usage, and outperforms EFE in predicting held-out data.